NSF PAR Search | NSF Public Access Repository

Note: When clicking on a Digital Object Identifier (DOI) number, you will be taken to an external site maintained by the publisher. Some full text articles may not yet be available without a charge during the embargo (administrative interval).
What is a DOI Number?

Some links on this page may take you to non-federal websites. Their policies may differ from this site.

DiaBlo: Diagonal Blocks Are Sufficient For Finetuning

Gurses, Selcuk; Zhang, Aozhong; Deng, Yanxia; Dong, Xun; Li, Xin; Wang, Naigang; Yin, Penghang; Yang, Zi (April 2026, The Fourteenth International Conference on Learning Representations)

Fine-tuning is a critical step for adapting large language models (LLMs) to domain specific downstream tasks. To mitigate the substantial computational and memory costs of full-model fine-tuning, Parameter-Efficient Fine-Tuning (PEFT) methods have been proposed to update only a small subset of model parameters. However, performance gaps between PEFT approaches and full-model fine-tuning still exist. In this work, we present DiaBlo, a simple yet effective PEFT approach that updates only the diagonal blocks of selected model weight matrices. Unlike Low-Rank Adaptation (LoRA) and its variants, DiaBlo eliminates the need for low-rank matrix products, thereby avoiding the reliance on auxiliary initialization schemes or customized optimization strategies to improve convergence. This design leads to stable and robust convergence while maintaining comparable memory efficiency and training speed to LoRA. Moreover, we provide theoretical guarantees showing that, under mild low-rank conditions, DiaBlo is more expressive than LoRA in the linear problem and converges to a stationary point of the general nonlinear full fine-tuning. Through extensive experiments across a range of tasks—including commonsense reasoning, arithmetic reasoning, code generation, and safety alignment—we show that fine-tuning only diagonal blocks is sufficient for strong and consistent performance. DiaBlo not only achieves competitive accuracy but also preserves high memory efficiency and fast fine-tuning speed. Codes are available at https://github.com/ziyangjoy/DiaBlo.
more » « less
Full Text Available
CLoQ: Enhancing Fine-Tuning of Quantized LLMs via Calibrated LoRA Initialization

Deng, Yanxia; Zhang, Aozhong; Gurses, Selcuk; Wang, Naigang; Yang, Zi; Yin, Penghang (August 2025, Transactions on machine learning research)

Fine-tuning large language models (LLMs) using low-rank adaptation (LoRA) has become a highly efficient approach for downstream tasks, particularly in scenarios with limited computational resources. However, applying LoRA techniques to quantized LLMs poses unique challenges due to the reduced representational precision of quantized weights. In this paper, we introduce CLoQ (Calibrated LoRA initialization for Quantized LLMs), a simplistic initialization strategy designed to overcome these challenges. Our approach focuses on minimizing the layer-wise discrepancy between the original LLM and its quantized counterpart with LoRA components during initialization. By leveraging a small calibration dataset, CLoQ quantizes a pre-trained LLM and determines the optimal LoRA components for each layer, ensuring a strong foundation for subsequent fine-tuning. A key contribution of this work is a novel theoretical result that enables the accurate and closed-form construction of these optimal LoRA components. We validate the efficacy of CLoQ across multiple tasks such as language generation, arithmetic reasoning, and commonsense reasoning, demonstrating that it consistently outperforms existing LoRA fine-tuning methods for quantized LLMs, especially at 2-bit.
more » « less
Full Text Available
Cloq: Enhancing fine-tuning of quantized llms via calibrated lora initialization

Deng, Yanxia; Zhang, Aozhong Zhang; Gurses, Selcuk; Wang, Naigang; Yang, Zi; Yin, Penghang (August 2025, Transactions on machine learning research)

Full Text Available
COMQ: A Backpropagation-Free Algorithm for Post-Training Quantization

https://doi.org/10.1109/ACCESS.2025.3576737

Zhang, Aozhong; Yang, Zi; Wang, Naigang; Qi, Yingyong; Xin, Jack; Li, Xin; Yin, Penghang (June 2025, IEEE Access)

Full Text Available
Diagonal Gaussian mixture models and higher order tensor decompositions

https://doi.org/10.3934/naco.2024053

Guo, Bingni; Nie, Jiawang; Yang, Zi (September 2025, Numerical Algebra, Control and Optimization)

Full Text Available
MagR: Weight Magnitude Reduction for Enhancing Post-Training Quantization

Zhang, Aozhong; Wang, Naigang; Deng, Yanxia; Li, Xin; Yang, Zi; Yin, Penghang (December 2024, Advances in Neural Information Processing Systems 2024)

Full Text Available
MagR: Weight Magnitude Reduction for Enhancing Post-Training Quantization

https://doi.org/10.52202/079017-2702

Deng, Yanxia; Li, Xin; Wang, Naigang; Yang, Zi; Yin, Penghang; Zhang, Aozhong (January 2024, Neural Information Processing Systems Foundation, Inc. (NeurIPS))

Full Text Available
The Multi-Objective Polynomial Optimization

https://doi.org/10.1287/moor.2023.0200

Nie, Jiawang; Yang, Zi (November 2024, Mathematics of Operations Research)

The multi-objective optimization is to optimize several objective functions over a common feasible set. Because the objectives usually do not share a common optimizer, people often consider (weakly) Pareto points. This paper studies multi-objective optimization problems that are given by polynomial functions. First, we study the geometry for (weakly) Pareto values and represent Pareto front as the boundary of a convex set. Linear scalarization problems (LSPs) and Chebyshev scalarization problems (CSPs) are typical approaches for getting (weakly) Pareto points. For LSPs, we show how to use tight relaxations to solve them and how to detect existence or nonexistence of proper weights. For CSPs, we show how to solve them by moment relaxations. Moreover, we show how to check whether a given point is a (weakly) Pareto point or not and how to detect existence or nonexistence of (weakly) Pareto points. We also study how to detect unboundedness of polynomial optimization, which is used to detect nonexistence of proper weights or (weakly) Pareto points. Funding: J. Nie is partially supported by the National Science Foundation [Grant DMS-2110780].
more » « less
Full Text Available
Dehomogenization for completely positive tensors

https://doi.org/10.3934/naco.2022037

Nie, Jiawang; Tang, Xindong; Yang, Zi; Zhong, Suhan (January 2023, Numerical Algebra, Control and Optimization)

Full Text Available
Theory and application of the vector pair correlation function for real-space crystallographic analysis of order/disorder correlations from STEM images

https://doi.org/10.1063/5.0058928

Funni, Stephen D.; Yang, Zi Jin; Cabral, Matthew J.; Ophus, Colin; Chen, Xiang M.; Dickey, Elizabeth C. (September 2021, APL Materials)

Full Text Available

« Prev Next »

Search for: All records